Back

npj Systems Biology and Applications

Springer Science and Business Media LLC

Preprints posted in the last 90 days, ranked by how well they match npj Systems Biology and Applications's content profile, based on 125 papers previously published here. The average preprint has a 0.09% match score for this journal, so anything above that is already an above-average fit.

1
Scalable biophysical constraints for physiologically consistent metabolic states

Toumpe, I.; Weilandt, D. R.; Narayanan, B.; Fengos, G.; Hatzimanikatis, V.; Miskovic, L.

2026-07-09 systems biology 10.64898/2026.07.03.736321 medRxiv
Top 0.1%
49.1%
Show abstract

Systems biology aims to develop predictive models that connect molecular mechanisms to cellular behavior. Genome-scale metabolic models are among the most widely used frameworks for integrating stoichiometric, thermodynamic, and omics-derived information to predict feasible metabolic phenotypes. However, cellular metabolism operates on timescales governed by enzyme kinetics and by the relationship between metabolic fluxes and metabolite pool sizes. In steady-state metabolic models, this relationship can be expressed in terms of metabolite turnover rates, defined as flux-to-pool-size ratios that quantify how rapidly metabolite pools are renewed. As a result, physiologically consistent steady-state solutions should not only satisfy mass-balance and thermodynamic constraints but also exhibit turnover rates consistent with enzyme-mediated cellular dynamics. Current constraint-based approaches can admit many steady-state flux-concentration states that do not account for turnover rates, resulting in phenotypes incompatible with realistic metabolic dynamics, even when multiple types of data are imposed. Here, we present METEOR-K, an optimization framework that links steady-state metabolic fluxes to metabolite concentrations via turnover rate constraints to identify dynamically plausible flux-concentration reference states. Because these constraints reshape the feasible solution space, we also introduce turnover-rate-aware sampling strategies to efficiently explore the resulting feasible region. We applied METEOR-K to models of increasing scope and scale, including a reduced glycolysis pathway, anaerobic E. coli, and near-genome-scale ovarian cancer models. METEOR-K narrowed the admissible steady-state solution space, reduced uncertainty in feasible flux-concentration states, and improved local dynamic behavior. In nonlinear ODE simulations of bioreactor cultivation and drug-response scenarios, METEOR-K-derived states produced intracellular response times compatible with growth-supporting metabolic operation and perturbation recovery. Overall, these results establish metabolite turnover rates as scalable biophysical constraints that improve the physiological consistency of steady-state metabolic modeling. Because turnover rates encode flux-to-pool-size timescale constraints, METEOR-K moves part of physiological-consistency assessment upstream of kinetic parameterization, yielding better-suited flux-concentration reference states for kinetic modeling and dynamic prediction.

2
MechAInistic: An LLM-guided Multi-Agent System for Reasoning over Genome-Scale Constraint-Based Metabolic Models

Loecker, J.; Pujara, N.; Bryant, W.; Puniya, B. L.; Packrisamy, P.; Hamed, A.; Helikar, T.

2026-05-13 systems biology 10.64898/2026.05.11.723319 medRxiv
Top 0.1%
39.5%
Show abstract

Constraint-based metabolic modeling is a powerful way to study the mechanistic basis of cellular states and disease, but effective use demands substantial computational expertise and careful coordination of multi-step analyses. We developed MechAInistic to lower this barrier enabling researchers to ask complex biological questions in natural language. MechAInistic is a multi-agent system harnessing large language models organized around an Architect-Reviewer pattern that that converts a natural-language question into an executable, model-grounded workflow and produces a structured report. It supports pathway comparison, perturbation analysis, drug-target exploration, and literature interpretation across healthy and disease paired states. We evaluated MechAInistics therapeutic hypothesis generation using two immune-cell use-cases. For rheumatoid arthritis/healthy Naive B models, it identified mitochondrial metabolic rewiring and nominated Devimistat/CPI-613 as an investigational OGDH-centered hypothesis. In CD4+ Th17 multiple sclerosis/healthy models, the workflow identified NADP-dependent isocitrate dehydrogenase as the optimal target and proposed Ivosidenib as an FDA-approved repurposing candidate. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=83 SRC="FIGDIR/small/723319v1_ufig1.gif" ALT="Figure 1"> View larger version (19K): org.highwire.dtl.DTLVardef@1b5c1d1org.highwire.dtl.DTLVardef@1c798cforg.highwire.dtl.DTLVardef@10161d3org.highwire.dtl.DTLVardef@1bd7dce_HPS_FORMAT_FIGEXP M_FIG C_FIG

3
Systems modeling identifies phenotype-determining signaling pathways controlled by phosphatase PTPRJ in diverse receptor tyrosine kinase activation settings

Hart, W. S.; Knight, K. M.; Rizzo, S.; Lee, S. H.; Fetter, R.; Thevenin, D.; Lazzara, M. J.

2026-05-04 systems biology 10.64898/2026.04.30.721884 medRxiv
Top 0.1%
30.2%
Show abstract

Protein tyrosine phosphatase receptor J (PTPRJ) restrains cell proliferation and migration by dephosphorylating receptor tyrosine kinases (RTKs) including the epidermal growth factor receptor (EGFR). PTPRJ is a purported tumor suppressor, and alterations to its expression and/or function are associated with colorectal, breast, lung, and other cancers. While there is interest in controlling PTPRJ-regulated phenotypes, efforts are limited by the complexity of PTPRJ-mediated signaling. PTPRJ dephosphorylates multiple RTKs, and the degree to which PTPRJ control of signaling and phenotypes depends on local cellular RTK activation profiles is unknown. To probe the context dependence of PTPRJ signaling regulation, we collected signaling measurements across 16 pathway nodes at two time points in a panel of HSC3 carcinoma cells engineered with different PTPRJ expression profiles. Cells were treated with three different RTK ligands, and paired phenotype measurements (viability, wound healing, xCELLigence cell index) were made. Partial least squares regression models were developed to predict relationships between PTPRJ-regulated signaling pathways and cell phenotypes. The model effectively separated contributions to variance arising from the PTPRJ expression background and growth factor context. In testing model predictions, we demonstrated that PTPRJ suppressed MET-induced cell cell proliferation via regulation of a HER3/AKT signaling axis that stabilized PTPRJ expression through an unanticipated feedback mechanism. We also found that PTPRJ regulated HSC3 cell migration via JNK signaling that was preferentially activated by MET. Our results identify new regulatory nodes through which PTPRJ influences cancer cell phenotypes and demonstrates that these processes preferentially occur in the context of distinct RTK activation states.

4
A new genome-scale model enables prediction of cancer metabolic dependencies

Dinh, H. V.; Zoitou, A.; Zhang, J.; Shen, Y.

2026-07-09 systems biology 10.64898/2026.06.30.735578 medRxiv
Top 0.1%
26.5%
Show abstract

Cancer cells rewire metabolism to support proliferation. Intriguingly, divergent metabolic choices are made to attain this common goal. Identifying the unique metabolic requirements for a specific cell has profound implications for cancer biology and precision medicine. Genome-scale metabolic models (GEMs) have emerged as powerful tools to systematically characterize, understand, and predict metabolism of cells and tissues. Despite being comprehensive, the current GEMs remain limited in their predictive power. Here, we present a new GEM of human cells, in silico Human Metabolic Essentiality (iHME), that significantly improves the prediction of metabolic dependencies at a reduced computational cost. Wse rationally downsized, curated, and corrected previous models to remove unsupported metabolic redundancies, which led to a slim model containing 4,377 reactions, 3,241 metabolites, and 1,825 genes. When used to reconstruct metabolic networks of 1,103 cancer cell lines, iHME recalled on average 84.6% of experimental essential genes, which is two-fold increase over previous models. Cholesterol biosynthesis was revealed to be the most reliably predicted pathway with alternative dependencies. Finally, we applied the model to reconstruct individualized networks and predict essential gene profiles for 8,384 patient tumor samples. Glucose transporter SLC2A1 (GLUT1) was identified as a context-specific dependency for head and neck cancers and ovarian cancer. Likewise, CDP-diacylglycerol synthase CDS2 was identified for skin cancer. Overall, iHME is a new genome-scale model for prediction of metabolic dependency at higher accuracy and computational efficiency.

5
Bidirectional network hubs: NT-genes as optimal targets for partial cancer reversal

Gil Perez, G. J.; Perez Rodriguez, R.; Gonzalez, A.

2026-04-30 systems biology 10.64898/2026.04.27.721122 medRxiv
Top 0.1%
22.8%
Show abstract

BackgroundThe complexity of gene regulatory networks, involving thousands of genes, poses a fundamental challenge to understanding cancer phenotype reversal. However, recent evidence suggests that the effective dimensionality of normal and tumor transcriptional manifolds is low, and that small panels of genes can discriminate perfectly between normal and tumor samples. MethodsWe build upon two previously developed concepts: (i) highly accurate normal and tumor gene markers (namely, N-and T-markers), defined as genes with exclusive expression intervals in normal and tumor samples, respectively; and (ii) gene deregulation networks (GDNs), represented as directed acyclic graphs encoding causal relationships between gene deregulation events. A subset of genes appearing in both marker classes (NT-markers) act as bridging nodes between the N-and T-GDNs. Starting from these elements, we introduce a quantitative dynamical model based on node frequency and connectivity to assess how gene intervention effects propagate through the GDN and thereby predict their overall impact on the tumor tissue. ResultsAccording to the model, interventions on pure T-markers (T-markers that are not NT-markers) produce effects largely confined to the T-GDN, with a minimal perturbation of the N-network. Interventions on pure N-markers (N-markers that are not NT-markers) generate a perturbation of both networks, but with limited effect. In contrast, interventions on NT-markers with high activation frequency in both tumors and normal state (e.g., AGER in lung adenocarcinoma: 98% in tumor samples, 75% in normal samples) can induce bidirectional phenotype shifts. For an effective combination of targets, coverage across tumor samples must be maximized. At the same time, in the T-GDN the number of nodes unreached by the reverse cascade following the intervention must be minimized, as these regions may act as escape routes for the tumor. Escape probability further depends on the tumor stage and the tumors activation rate of new T-genes. When targeting NT genes, high frequency in normal samples should also be prioritized. ConclusionsHigh-frequency NT-genes, due to dual network connectivity and tissue relevance, represent optimal targets for achieving at least partial phenotype reversal. This framework provides a quantitative guide for prioritizing gene therapy targets and designing combination strategies that balance coverage, escape minimization, and normal tissue relevance.

6
Mechanistically informed adaptive dosing for cancer immunotherapy using AI-guided decision making

Garg, A.; Das, S. S.; Sivadasan, N.; Roy, A.; Chakrabarty, B.

2026-07-08 systems biology 10.64898/2026.06.09.730783 medRxiv
Top 0.1%
21.6%
Show abstract

Optimizing dose and schedule remains a central challenge in oncology drug development, particularly for immunotherapies where fixed dosing regimens often fail to account for patient specific heterogeneity in tumor-immune dynamics. Here, we present a hybrid quantitative systems pharmacology-reinforcement learning-Monte Carlo Tree Search (QSP-RL-MCTS) framework for personalized immunotherapy dosing that formulates dose selection as a sequential decision-making problem. The approach integrates a mechanistic QSP model of prostate cancer immunotherapy, transcriptomics informed virtual patient populations and data driven AI system comprising reinforcement learning and Monte Carlo tree search. Reinforcement learning is used to learn adaptive generalized dosing policies that optimize treatment outcomes across the population, while Monte Carlo Tree Search provides forward-looking evaluation of RL predicted dosing trajectories to refine patient-specific decisions. On benchmarking against fixed dosing regimens of ipilimumab, the remission rate of the proposed model (95.2%) was comparable to the highest fixed dosing regimen of 10 mg/kg per dose while the median total dose (72 mg/kg) of the proposed model designed regimen was comparable to the lowest fixed dosing regimen of 3 mg/kg per dose. The model is generalizable across different dosing protocols and can be extended to predict optimal dose under different therapeutic scenarios. Analysis of the learned dosing trajectories enables stratification of patients into distinct response groups and identifies drug activity rate as the dominant determinant of long-term treatment outcome. These results demonstrate how mechanistically guided artificial intelligence can transform population-level dose optimization into patient-specific, biologically interpretable treatment strategies for precision immuno-oncology.

7
Growth-resolved genome-scale metabolic modeling of Priestia megaterium SR7 validated by chemostat and 13-C flux analysis

Chang, K. Y. W.; Song, Y.; Hing, N. Y. K.; Vethathirri, R. S.; Wang, Y.; Thompson, J. R.

2026-05-30 systems biology 10.64898/2026.05.27.728139 medRxiv
Top 0.1%
19.0%
Show abstract

Priestia megaterium SR7 is a promising candidate chassis for bioprocess engineering, but its development is limited by the availability of condition-grounded, mechanistic models that can translate experimental measurements into predictive design hypotheses. Here, we present PMSR7, a genome-scale metabolic model for SR7, and evaluate it under a growth-resolved chemostat framework spanning a dilution-rate series. Stable steady states were established across the growth regime, with the highest dilution rate (D = 1.1538 h-{superscript 1}) excluded from growth interpretation due to biomass collapse. Extracellular carbon fluxes were quantified by NMR and reported as mean {+/-} SD, providing an experimental basis for model comparison. PMSR7 was benchmarked using MEMOTE against representative reference reconstructions, supporting structural consistency suitable for constraint-based analyses. Under growth-resolved simulations, ATP demand scaled linearly with growth rate, enabling inference of maintenance-energy behavior across the regime. Growth-dependent feasibility and magnitude of overflow secretion were evaluated for acetate, lactate, and formate using feasible-space analyses, highlighting both agreement and regime-sensitive limitations. Finally, growth-resolved leucine and valine production was assessed in both raw and fold-change space, with experimental means compared against median-based summaries of sampled model distributions to account for feasible-space skew. Together, these results establish PMSR7 as a reproducible, quality-benchmarked platform for SR7 chassis development and provide a framework for iterative experimental integration in non-model organisms, where the dominant challenge is achieving congruence between measured physiology and model-feasible behavior.

8
Intricate Dynamical Cross-Talk Between p53 Protein and Cell Cycle Regulators Governs Mammalian Cell Fate

Charan, K.; Kar, S.

2026-06-10 systems biology 10.64898/2026.06.07.730771 medRxiv
Top 0.1%
18.2%
Show abstract

In mammalian cells, under normal circumstances, the p53 protein exhibits oscillatory dynamics in response to DNA damage and maintains the cells in a cell-cycle-arrested state. Intriguingly, some cells can escape this cell-cycle-arrested state even after prolonged DNA damage, and often undergo mitotic catastrophe. In this context, the precise role of p53 dynamics and its complex interplay with cell-cycle regulation remain poorly understood. Herein, by constructing a comprehensive network model, we have identified crucial crosstalk regulations between the p53 protein and key cell-cycle regulators that enable some cells to escape cell-cycle arrest during prolonged DNA damage. The model further illustrates a probable cellular mechanism underlying mitotic catastrophe and predicts ways to induce it in a therapeutically relevant manner.

9
Talk2QSP: Deriving Executable Scenarios from Unstructured Literature via Human-in-the-Loop Agents

Kazemeini, A.; Prieto, J.; Balaji Kuttae, S.; Siokis, A.; Singh, G.; Passban, P.; Andreani, T.

2026-05-11 systems biology 10.64898/2026.05.06.723244 medRxiv
Top 0.1%
18.1%
Show abstract

Quantitative Systems Pharmacology (QSP) models play an inherently interventional role in pharmaceutical research and development, functioning as executable causal systems for designing, evaluating, and replacing clinical trials. However, deploying QSP as an experimental planning engine remains constrained by the difficulty of translating unstructured literature descriptions of clinical or preclinical scenarios into reproducible, simulation-ready model interventions. Motivated by this issue, we propose an agent-based framework that operationalizes QSP models as intervention-ready experimental systems by automatically extracting and executing literature-derived scenarios. The framework combines semantic grounding of model entities with a large language model (LLM)-driven Scenario Extractor and a dual-agent Scenario Mapper. Rather than relying on opaque, single-shot reasoning, our pipeline converts free-text interventions into precise parameter configurations through discrete, verifiable work orders. Moreover, our dynamic Human-in-the-Loop (HITL) strategy empowers modelers to resolve biological ambiguities interactively. Across four diverse kinetic ordinary differential equation (ODE)/QSP models and seven Subject Matter Expert (SME)-curated literature scenarios, our model resolved all selected scenarios into correct executable parameter changes, including multi-dose interventions, unit conversions, no-op scenarios, and ambiguity-triggered HITL cases, demonstrating that structured collaboration between experts and agentic systems can resolve scenarios that standalone raw Systems Biology Markup Language (SBML) reasoning LLM calls handle unreliably.

10
A Novel Network Approach to Identify Sample-Specific Context-Informed Metabolic Signatures During Developmental Processes

Lee, E.; Koppayi, A.; Veiga-Lopez, A.; Penalver Bernabe, B.

2026-05-22 bioinformatics 10.64898/2026.05.20.726642 medRxiv
Top 0.1%
15.4%
Show abstract

Metabolism plays an essential role in cellular processes: development, growth, differentiation, and determination of cell identity. Understanding how metabolic processes dynamically change across cell types, stages, and environmental conditions is crucial for studying developmental biology, aging, and disease progression. Genome-wide metabolic models (GEMs) are a powerful network-based tool for studying these processes by integrating omics data to model context-specific metabolism. However, current approaches, such as Flux Balance Analysis (FBA), have limitations in addressing the dynamic nature of metabolism across developmental stages at a sample-specific resolution. To address this, we introduce a novel network-based method for analyzing cell and stage specific metabolic flow using directed and weighted metabolic networks that account for sample-specific transcriptomic data. We apply this method to study ovarian follicle development, providing a deeper understanding of intra-cellular metabolic processes, identifying key metabolites, enzymes, and potential markers for follicular maturation, important for IVF. By incorporating biologically meaningful data, this approach bridges the gap between theoretical metabolic network models (GEMs) and experimental observations, offering a systems-level view of metabolic dynamics in developmental and understudied contexts.

11
Weak form Scientific Machine Learning for Systems Biology: A Tutorial on WENDy

Heitzman-Breen, N.; Lyons, R.; Jain, P.; Jolly, M. K.; Bortz, D. M.

2026-07-09 systems biology 10.64898/2026.07.02.735880 medRxiv
Top 0.1%
15.3%
Show abstract

Mechanistic ordinary differential equation models are widely used in systems biology to represent biochemical networks, population dynamics, cell-state transitions, and other biological processes; however, their predictive value depends critically on accurate parameter estimation from noisy and often sparse experimental data. In this tutorial, we present the Weak-form Estimation of Nonlinear Dynamics (WENDy) method as a forward-solver-free approach that reformulates parameter estimation as a covariance-corrected weak-form regression problem by integrating the model equations against compactly supported test functions. We present the background on the methodology through the lens of the familiar logistic equation, and we demonstrate applications of the method on real experimental data through two systems biology examples: a glycolytic oscillator with relatively dense time-course data and a sparse epithelial-mesenchymal cellstate transition model with multiple experimental replicates. Ultimately, using WENDy, we estimate interpretable biological parameters with uncertainty for systems with noisy and sometimes sparse available experimental data.

12
Human Genome-Scale Models of Metabolism and Gene Expression Reveal Resource Constraints of Cancer Cell Lines

Baghdassarian, H. M.; Di Giusto, P.; Tibocha-Bonilla, J.; Armingol, E.; Gopalakrishnan, S.; Dworkin, L.; Yang, L. M.; Lewis, N. E.

2026-06-03 systems biology 10.64898/2026.05.30.728988 medRxiv
Top 0.1%
14.8%
Show abstract

Genome-scale metabolic models (M-models) provide mechanistic insight into intracellular metabolism by simulating fluxes subject to nutrient and energy resource constraints. However, they cannot account for a major component of resource allocation, since they do not explicitly account for the cost of producing and maintaining enzymes. Genome-scale models of metabolism and gene expression (ME-Models) address this by including gene expression reactions, but these have only been developed for prokaryotes due to the additional complexity and challenges of modeling eukaryotes. Here, we present the human ME-Model, which encodes transcription, translation, complex formation, and turnover reactions for all enzymes catalyzing metabolic reactions, and couples these processes to constrain metabolic fluxes. We introduce humanME, a Python package to build and analyze human ME-Models. With this, we constructed 16 cancer cell line ME-Models. We found that resource constraints improve growth-rate predictions, and that ME-Model flux predictions are more biologically plausible and efficient. Moreover, transcriptional fluxes recapitulate RNA-Seq expression levels, with discrepancies revealing potential trade-offs involving multiple cellular objectives. Finally, the ME-Model recapitulates the Warburg effect, with increasing growth rate inducing glycolytic shifts, in part due to machinery costs of the electron transport chain. Altogether, we show ME-modeling can mechanistically link gene expression, resource allocation, and metabolism in human cells, substantially expanding the predictive scope of constraint-based models.

13
Gene Regulatory Networks that support Multi-Fate Cellular Decisions

BV, H.; Adigwe, S.; Jolly, M. K.; Gedeon, T.

2026-07-15 systems biology 10.64898/2026.07.13.738161 medRxiv
Top 0.1%
14.7%
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWCell fate decisions are driven by gene regulatory networks (GRNs). While the mutually inhibitory toggle switch effectively models binary fate decisions, fully connected inhibitory networks with more than two nodes fail to capture multi-fate decisions due to the low prevalence of "single high states", where only a single master regulator is highly expressed. The goal of this study is to find network structures that support all single high states. We find that the only network that attains the highest possible prevalence of all single high states within the set of monotone Boolean (MB) models is completely disconnected. Since biological networks typically require connectivity, we investigate network structures that support equipotency, where all single high states have equal prevalence within MB models. Finally, we characterize the networks that support multistability between all single high states, finding that it is possible only in networks in which each node either has self-activations or is inhibited by every other network node. Our findings provide a theoretical framework for understanding the network design principles that can support simultaneous differentiation into multiple distinct cell types.

14
The first digital twin of Enterococcus faecium metabolism reproduces high-throughput phenotyping data

Rasmi, D. S.; Krishnan, J.; Hashem, Y. A.; Palsson, B.; Khashef, M. T.; Monk, J.; Aziz, R. K.

2026-05-06 systems biology 10.64898/2026.05.01.720924 medRxiv
Top 0.1%
14.5%
Show abstract

Enterococci are Gram-positive opportunistic pathogens responsible for a wide range of nosocomial infections. One enterococcocal species, Enterococcus faecium, is steadily increasing in prevalence and has been listed among major multidrug-resistant ESKAPE pathogens. To gain systems-level insights into its metabolism and support discovery of potential therapeutic targets, we constructed iDR479, a comprehensive manually curated genome-scale metabolic model (GEM) to serve as a digital twin for E. faecium TX0016 (strain DO). The reconstruction was curated through extensive homology searches and literature evidence, and further refined and gap-filled through experimental validation. Phenotypic profiling using Biolog microarrays enabled assessment of carbon source utilization, while amino acid leave-out growth assays allowed the evaluation of auxotrophies. The final refined model is 100% accurate in predicting amino acid auxotrophy and 85% accurate in predicting growth on sole carbon sources. Discrepancies between model predictions and experimental phenotypes identified specific knowledge gaps across metabolic pathways, including unresolved carbon source utilization phenotypes, e.g., psicose, sorbitol, and palatinose utilization. Those gaps will guided future experimental characterization. Additionally, gene essentiality analysis was conducted to evaluate the predictive capacity of iDR479 model. Since no experimental gene essentiality data are currently available for E. faecium, model predictions were compared against Tn-seq experimental results from E. faecalis MMH594. Under simulated rich medium conditions, iDR479 achieved 86.7% concordance with the experimental essentiality results of E. faecalis MMH594. iDR479 thus provides a framework for studying E. faecium, offers insights into its metabolic network, and serves as a source for guiding future research and identification of therapeutic targets.

15
DrugPTM-Bench: A Large-Scale Dataset for Predictive Modeling of Drug-Induced Cell Type-Specific Protein Post-Translational Modifications

Badkul, A.; Mottaqi, M.; Xie, L.; Xie, L.

2026-04-30 systems biology 10.64898/2026.04.27.721113 medRxiv
Top 0.1%
12.9%
Show abstract

Protein post-translational modifications (PTMs), particularly phosphorylation, serve as the primary "molecular switches" that orchestrate cellular signaling and drug response. While PTM dysregulation is a hallmark of cancer and neurodegeneration, the lack of standardized, drug-perturbed datasets has hindered the development of predictive models capable of capturing context-dependent PTM responses. Effective predictive modeling must therefore integrate multidimensional data, including the specific drug, dosage, treatment duration, cellular background, and the modified site. However, existing PTM resources remain largely static and fail to capture drug-induced regulation across these critical dimensions. To address this gap, we present DrugPTM-Bench, a curated, large-scale benchmark derived from decryptM-derived dose-dependent PTM measurements, standardizing site-level drug response across 7 cancer cell lines, 27 drugs, and 11,167 proteins. Comprising 99.5% phosphorylation events, the dataset includes six time points, 16 dosage levels, and pEC50 potency values (half-maximal effective concentration). We formulate a classification task to identify upregulated, downregulated, or unchanged PTM sites (following a drug treatment), a critical step in deciphering drug Mechanism of Action (MoA) and target engagement. Our evaluation reveals that in protein-disjoint out-of-distribution (OOD) setting, baseline machine learning and deep learning models struggle to recover minority regulation classes, while standard rebalancing strategies improve recall only at the cost of precision and overall F1-score. These results indicate that current methods do not learn robust decision boundaries between regulated and unchanged PTM events. DrugPTM-Bench provides a phosphoproteomics benchmark for modeling drug-induced PTM regulation in imbalanced biological settings. Beyond classification, DrugPTM-Benchs retention of pEC50 values, drug perturbation profiles, and site-level sequence context enables additional predictive tasks including drug potency regression, mechanism-of-action prediction from PTM fingerprints, and drug-specific PTM site sensitivity ranking, establishing a multi-task benchmark for PTM-centric drug discovery. Ultimately, DrugPTM-Bench establishes a rigorous framework for developing robust, context-aware models to elucidate drug MoA and signaling dynamics.

16
SPARK: A Systems-level Computational Framework for Reconstructing Transcriptomic State Organisation in Lung Adenocarcinoma

Kulkarni, R.; Sengupta, A.; Kumar, R.

2026-06-11 bioinformatics 10.64898/2026.06.08.730929 medRxiv
Top 0.1%
12.7%
Show abstract

Lung adenocarcinoma (LUAD) exhibits substantial molecular heterogeneity, which complicates tumour stratification and limits the ability of mutation-centric models to capture tumour behaviour and predict patient outcomes. This study investigates whether coordinated transcriptomic programs can provide a systems-level representation of tumour states. Bulk RNA-sequencing data from the TCGA-LUAD cohort were analysed to reconstruct pathway-level transcriptomic organisation using a stability-optimised network framework (SPARK). This analysis identified eight transcriptomic modules representing coordinated biological processes active across tumours. Module activity scores were subsequently used to derive a composite Transcriptomic Risk Score through elastic-net Cox proportional hazards modelling. The resulting risk score showed a significant association with overall survival in the discovery cohort and improved prognostic discrimination beyond clinical variables. An independent evaluation in the CPTAC-LUAD cohort confirmed the prognostic signal and preserved risk stratification across patient groups. Unsupervised clustering of module activity further revealed three transcriptomic patient groups characterised by distinct biological programs, genomic alteration patterns, and survival outcomes. Single-cell analysis also demonstrated that the identified transcriptomic modules reflect coordinated organisation of the tumour-immune-stromal ecosystem across cellular compartments. Together, these findings suggest that LUAD heterogeneity can be organised into coordinated transcriptomic programs with measurable clinical relevance, providing a systems-level framework for representing tumour molecular states.

17
Gene Regulatory Network Inference reveals tcf4 as a key a player in neuroblastoma gene expression circuitry

Koering, C.; Vallin, E.; Picard, F.; Gonin-Giraud, S.; Gandrillon, O.

2026-07-08 cancer biology 10.64898/2026.07.07.737136 medRxiv
Top 0.1%
12.6%
Show abstract

Neuroblastoma (NB), a pediatric cancer arising from disrupted sympathetic neuron differentiation, exhibits marked heterogeneity and limited therapeutic options. To better understand its molecular circuitry dynamics, we applied CardamomOT, a novel Gene Regulatory Network (GRN) inference framework, to single-cell RNA-seq data from patient-derived tumoroids. This approach models gene regulation via piecewise deterministic Markov processes, capturing transcriptional bursting and protein-mediated feedback, overcoming limitations of RNA velocity (e.g., gene independence and lack of biological time). We identified a continuous chromaffin-to-sympathoblast differentiation trajectory along which we selected 85 dynamically relevant genes enriched in cell cycle and DNA replication functions. Notably, 9 genes overlapped with those driving normal sympathoadrenal differentiation, underscoring tumor-normal tissue similarity. The inferred 85-genes network reproduced quite well experimental gene expression patterns in silico, and allowed to predict protein-level dynamics. Furthermore, it allowed to predict the effect of perturbations (both knock-out and overexpression) of hub genes (e.g., tcf4 and PLK1). We show that those perturbations significantly altered cell fate proportions in silico, with tcf4 KO increasing chromaffin-like cells and reducing proliferative late sympathoblasts. Predictions regarding tcf4 were tested using drug inhibition as a proxy for the gene KO. Using the BET inhibitor JQ1 indeed induced profound effect on the transcriptomic identity of our tumoroids. All of the 50 predicted tcf4 target genes were found to be significantly altered by JQ1 treatment. Finally cell fate proportions were also altered ex vivo closely resembling the predicted output. Our work therefore demonstrates that NB tumoroids retain a dynamic, differentiation-like architecture amenable to GRN modeling. Predicted druggable targets offer testable therapeutic avenues, including repurposing BET inhibitors or PLK1 inhibitors, potentially in combination.

18
Mechanistic pathway modeling reveals how IL-10 generates pleiotropic immune responses

Marti Baena, Q.; Segura-Morales, C.; Garcia Ojalvo, J.; Serrano, L.

2026-06-07 systems biology 10.64898/2026.06.02.729479 medRxiv
Top 0.1%
12.4%
Show abstract

IL-10 is a key anti-inflammatory cytokine whose activity is impaired in autoimmune diseases. However, IL-10 also promotes inflammation under certain conditions, limiting the efficacy of IL-10-based therapies. Because the principles underlying these opposing effects remain unclear, we developed a mathematical model of the IL-10 signaling pathway to understand how such pleiotropic responses arise. Considering that STAT3 signaling is buffered against changes in IL-10RB receptor affinity, we provide a minimal mechanistic explanation for IL-10 variants with anti-inflammatory or pro-inflammatory biased responses produced by an altered receptor affinity. By linking model-predicted pSTAT1 and pSTAT3 abundances with transcriptomic changes, we identified IL-10-responsive genes that are regulated at different pSTAT thresholds. This could explain how IL-10 elicits distinct downstream responses at different signaling strengths, and suggests that, depending on the required response, the affinity of IL10 for its receptors should not always be enhanced. Overall, our study highlights how modeling can help disentangle IL-10 pleiotropy to support the rational development of more effective IL-10-based therapies.

19
An extension of Modular Response Analysis for global perturbations and robust connectivity inference of gene regulatory networks.

Jimenez-Dominguez, G.; Audit, B.; Borgnat, P.; Ravel, P.; Arbona, J.-M.

2026-06-05 systems biology 10.64898/2026.06.02.729263 medRxiv
Top 0.1%
12.1%
Show abstract

Understanding how gene regulatory networks respond to global cell perturbations remains a central challenge in systems biology and network inference. Modular Response Analysis (MRA) provides a mathematical framework to infer gene-to-gene directed connectivity graphs from perturbation experiments; however, classical MRA captures direct gene-to-gene influences, and does not explicitly account for global stimuli that simultaneously change the graph. Here, we introduce MRA+, an extension of MRA, that incorporates the effect of global perturbations into gene-to-gene graph inference. MRA+ assumes a sequential experimental design in which targeted gene perturbations are followed by the application of a global stimulus, enabling the separation of connectivity changes from direct gene induction. The method estimates network connectivity under induced conditions and quantifies gene-specific induction strengths, which represent contributions to expression changes arising from mechanisms external to the inferred network. In the case of single-cell expression data, we present a bootstrap strategy to assess the robustness of inferred connectivity coefficients and propose a complementary criterion based on sign stability to interpret weak or non-significant estimates. Together, these developments provide a general framework for robust inference of gene connectivity graphs in the presence of global perturbations, applicable to diverse biological and experimental contexts.

20
Network-Level Characterization of Spontaneous Calcium Activity in an In-Vitro Alzheimer's Disease Model

Emenheiser, A. M.; Gentry, E.; Xue, H.; Alvarez, P.; O'Neill, K.; Cao, K.; Losert, W.

2026-06-01 biophysics 10.64898/2026.05.28.728474 medRxiv
Top 0.1%
11.9%
Show abstract

The neurodegenerative disorder Alzheimers disease (AD) is widely known for biomarkers such as amyloid beta plaques and tauopathy, as well as functional differences in memory and cognitive ability. Despite this devastating functional impact, a large body of work only focuses on molecular biomarkers of AD. In this study, we investigate collective neural dynamics in vitro and assess how network-level properties differ between a well-established model of familial AD (FAD) and a newly developed in vitro accelerated model (acAD). The new model system reliably develops the key structural characteristics of AD in three weeks, but its calcium dynamics had not been characterized previously. Spontaneous network dynamics influences information processing as part of the internal network state. Here we measure this spontaneous activity of a network of hundreds of cells in each field of view. We find that the FAD model has a larger fraction of hyperactive cells, while the acAD model displays similar characteristics to healthy cells. Additionally, the FAD model has altered cooperation between cells, losing a proportion of highly correlated cellular activities, both for fast and slow coupling among cells. The acAD model is again consistent with healthy networks. Since the acAD model does not show the same spontaneous network dysfunction seen in FAD, it can enable measurements of changes in learning and memory associated with the plasticity, rather than the structure of the internal network state.